Back

Pharmacoepidemiology and Drug Safety

Wiley

Preprints posted in the last 90 days, ranked by how well they match Pharmacoepidemiology and Drug Safety's content profile, based on 18 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
GLP Medications and Severe Post-COVID-19 Outcomes Among Individuals with Type 2 Diabetes Mellitus

Butzin-Dozier, Z.; Wang, L.-C.; Ji, Y.; Kumar, M.; Anzalone, A. J.; Hurwitz, E.; Patel, R. C.; Budhihartanto, A.; Buse, J. B.; Johnson, S.; Reusch, J.; Bramante, C.; Wong, R.; on behalf of the National Clinical Cohort Collaborative,

2026-07-06 epidemiology 10.64898/2026.07.03.26357246 medRxiv
Top 0.1%
19.0%
Show abstract

Background: Glucagon-like peptide-1 receptor agonist-based therapies (GLP) have recently emerged as promising treatments across a wide range of health conditions. These medications may have protective effects against severe long-term consequences of COVID-19 by promoting weight loss, exerting antihyperglycemic and anti-inflammatory effects, and providing cardiovascular and endothelial protection. Methods: We evaluated electronic health record data from a retrospective cohort of individuals in the National Clinical Cohort Collaborative. We included individuals with type 2 diabetes mellitus and comorbid COVID-19 who were prescribed either GLP (treatment) or a sodium-glucose co-transporter 2 inhibitor (SGLT2i) and subsequently developed acute COVID-19 between October 1, 2021, and April 1, 2023. We compared the 12-month cumulative incidence of mortality and Long COVID (Long COVID diagnosis and probable Long COVID via computational phenotype) between groups. We applied targeted maximum likelihood estimation to compare outcome risks by exposure status, controlling for covariates of interest. Results: We analyzed data from 14,215 individuals with COVID-19 and comorbid type 2 diabetes (mean age, 60 years; mean BMI, 37). Compared to SGLT2i, a prescription for GLP medication was associated with a lower risk of mortality (adjusted risk ratio [aRR] 0.71; 95% CI 0.53, 0.95), but not Long COVID diagnosis (aRR 1.01; 95% CI 0.80, 1.27) or probable Long COVID (aRR 0.94; 95% CI 0.88, 1.01). Conclusions: We found that among individuals with type 2 diabetes and comorbid COVID-19, a prescription for GLP vs. SGLT2i medications was associated with a lower risk of mortality, but not Long COVID.

2
Metformin and Severe Post-COVID-19 Outcomes Among Individuals with Diabetes Mellitus

Butzin-Dozier, Z.; Ji, Y.; Wang, L.-C.; Anzalone, A. J.; Olawore, O.; Hafen, R.; Hurwitz, E.; Kumar, M.; Patel, R. C.; Budhihartanto, A.; van der Laan, M.; Colford, J. M.; Hubbard, A. E.; Buse, J. B.; Johnson, S.; Reusch, J.; Chan, L. E.; Moffitt, R.; Wong, R.; Bramante, C.; on behalf of the National Clinical Cohort Collaborative,

2026-07-09 epidemiology 10.64898/2026.07.06.26357398 medRxiv
Top 0.1%
13.0%
Show abstract

Background: Metformin is one of the most commonly prescribed medications for individuals with diabetes and may provide protection against long-term sequelae of COVID-19. Methods: We evaluated a retrospective cohort of individuals in the National Clinical Cohort Collaborative with type 2 diabetes mellitus and COVID-19 who were prescribed metformin or a dipeptidyl peptidase-4 inhibitor (DPP4i) at least 30 days before the onset of acute COVID-19 between October 1, 2021, and November 15, 2023. We compared the 12-month cumulative incidence of Long COVID diagnosis (ICD-10 U09.9: Post COVID-19 condition, unspecified), probable Long COVID (based on a model-derived phenotype), and mortality between individuals prescribed metformin vs. DPP4i. We applied Super Learner and targeted maximum likelihood estimation to obtain risk ratios while adjusting for covariates of interest. Results: In our sample of 53,332 individuals with type 2 diabetes and COVID-19, we found that metformin prescription was associated with a lower risk of all-cause mortality after COVID-19 (adjusted risk ratio [aRR] 0.61, 95% CI 0.51, 0.73). We also observed that metformin users, compared to DPP4i users, had a slightly lower risk of probable Long COVID (aRR 0.87, 95% CI 0.81, 0.94) but did not detect a significant relationship with Long COVID diagnosis (aRR 0.90, 95% CI 0.68, 1.20), although we observed similar point estimates across Long COVID outcomes. Conclusions: These findings support the hypothesis that metformin prescription during acute COVID-19 may be associated with lower mortality among adults with diabetes. These analyses also provide modest evidence of a protective association against Long COVID in adults with diabetes, although estimates were imprecise.

3
Care Delivery Gap framework: a proof-of-concept patient-reported measure of guideline-referenced care-process omissions in sickle cell disease

Agbalalah, T.; Rowaiye, A.

2026-06-16 hematology 10.64898/2026.06.08.26355133 medRxiv
Top 0.1%
12.6%
Show abstract

Abstract Background:Sickle cell disease (SCD) is concentrated in sub-Saharan Africa, where delivery of guideline-referenced care remains challenging. Current evaluation approaches rely largely on access indicators and clinical outcomes, which do not directly measure care delivery. We developed the Care Delivery Gap (CDG) framework, a patient-reported approach for identifying care-process omissions, and conducted a proof-of-concept study to assess feasibility and explore variation across income strata. Methods: We conducted a cross-sectional framework-development study involving a proof-of-concept sample of 52 individuals with SCD or caregivers recruited through clinics and moderated SCD communities across Africa, North America, and Europe between June 2025 and March 2026. The CDG framework assessed patient-reported omissions in specialist involvement, follow-up continuity, cardiovascular screening, and biochemical surveillance. Analyses were descriptive. Results: Substantial multi-domain care-process omissions were identified despite high reported healthcare engagement. Across geographic income strata, cardiovascular screening was reported by 4/35 (11%) LMIC versus 16/17 (94%) HIC participants, and regular follow-up within the preceding 12 months by 14/35 (40%) versus 16/17 (94%), respectively. High CDG scores, representing 1 omissions across three or four domains, occurred in 20/35 (57%) LMIC compared with 1/17 (6%) HIC participants. Similar disparities were observed across specialist review and vitamin B12 surveillance domains. Conclusion: A structured patient-reported framework identified multi-domain omissions in guideline-referenced SCD care, including among individuals reporting healthcare access. The divergence between access indicators and reported care delivery suggests that service contact alone may not reflect care quality. The framework provides a feasible foundation for future process-level quality measurement in high-burden settings.

4
Integrating Causal Inference into Pharmacovigilance: Target Trial Emulations for Proactive Signal Detection of Atorvastatin Initiation in Medicare Beneficiaries

Rowan, C. G.; Tran, M.; Srivastava, S.

2026-07-10 epidemiology 10.64898/2026.07.01.26356874 medRxiv
Top 0.1%
11.8%
Show abstract

Importance: Adverse drug events in older adults are a substantial public health burden, yet spontaneous reporting systems detect them poorly owing to underreporting and the lack of a defined population. These limitations are of particular concern for older adults, who are underrepresented in pre-approval trials yet at elevated risk owing to polypharmacy, multimorbidity, and age-related changes in drug metabolism. Objective: To develop and apply an active, claims-based pharmacovigilance framework using sequential target trial emulation to detect adverse drug event signals in older adults, with atorvastatin as the initial application. Methods: Using Medicare fee-for-service claims (2017-2019), we studied statin-naive beneficiaries aged 65 years or older following myocardial or cerebral infarction. We emulated up to 14 daily sequential trials from the discharge date, classifying patients as initiating atorvastatin (A1), initiating a different medication (A2), or no new medication (A0); the primary contrast was A1 versus A2. For each trial, incident outcomes were ascertained and classified into 552 outcomes based on the Clinical Classifications Software Refined categories. Per-protocol effects were estimated over a 6-month follow-up period using Fine-Gray regression models weighted by the inverse probability of treatment and censoring, treating death as a competing risk, with the false discovery rate controlled via the Benjamini-Hochberg procedure. A signal was declared when the q-value was 0.10 or lower and the subdistribution hazard ratio (sHR) was 1.20 or greater in any prespecified analytic stratum (sensitivity analyses used thresholds of q 0.20 or lower and sHR 1.20 or greater). Results: Of 70,130 eligible patients, 39,948 initiated atorvastatin (A1) and 19,182 initiated another new medication (A2); after weighting, baseline characteristics were closely balanced. After excluding outcomes with sparse cell counts, 295 outcomes were analyzed; five met the primary signal detection criteria: valve disorders (sHR 1.71, 1.20 to 2.43); sprains and strains (sHR 1.79, 1.26 to 2.54); general sensation/perception symptoms (sHR 1.23, 95 percent CI 1.11 to 1.36); abnormal findings without diagnosis (sHR 1.55, 1.18 to 2.05); and prediabetes (sHR 1.71, 1.24 to 2.36). In the sensitivity analysis, we additionally detected posthemorrhagic anemia, hemorrhagic stroke, varicose veins, and other circulatory and skin conditions. Conclusions: An active, claims-based framework using sequential target trial emulation detected both expected and previously unrecognized adverse drug event signals following atorvastatin initiation in older adults, offering a systematic alternative to passive surveillance that can be extended to other commonly prescribed medications.

5
Heterogeneity of Treatment Effect of Aspirin and Clinically Significant Bleeding in Older Adults

Tzimas, G.; Tchoua, R. B.; Vanghelof, J. C.; Wolfe, R. C.; Cloud, G.; Mahady, S.; Du, L.; Ernst, M. E.; Wood, E. M.; Raicu, D. S.; Ket, S.; Shah, R. C.

2026-06-12 hematology 10.64898/2026.06.10.26355385 medRxiv
Top 0.1%
11.8%
Show abstract

Aim: The global population of older adults is growing, and older age is linked to higher bleeding risk. Although guidelines discourage aspirin for primary prevention in healthy older adults due to bleeding harms outweighing benefits, many continue taking it without a clear indication. It remains unclear whether all older adults face uniform aspirin-related bleeding risk or if certain subgroups are more vulnerable. Methods: We analyzed data from 19,114 ASPREE trial participants to develop machine learning models using 116 baseline variables. Random forest (RF) and random survival forest (RSF) models predicted 5-year bleeding risk, and participants were stratified into low, intermediate, and high-risk groups based on the 20th and 80th percentiles of predicted risk. We assessed heterogeneity of treatment effect (HTE) by testing treatment-by-risk group interactions on the relative scale using Fine-Gray models, and on the absolute scale using observed 5-year cumulative incidence rates. Results: Over a median follow-up of 4.7 years, 626 major bleeding events occurred. The RF model had moderate discrimination (AUC = 0.65, 95% CI: 0.63-0.67) and good calibration (Brier = 0.032, 95% CI: 0.029-0.034). Statistically significant HTE was observed on the relative scale, with the greatest relative increase in bleeding risk seen in the low-risk group (subdistribution hazard ratio = 2.26, 95% CI: 1.27-4.01). On the absolute scale, low-risk participants experienced higher bleeding with aspirin (absolute risk difference (ARD) = 1.17%, 95% CI: 0.37-1.95), but heterogeneity in ARDs was not statistically significant (Cochran's Q p > 0.45). Similar findings were observed when using the RSF model. Conclusion: Participants at lowest baseline bleeding risk experienced the greatest relative increase in bleeding risk with aspirin therapy. We found statistically significant heterogeneity in treatment effects on the relative but not absolute scale. These findings support an individualized, risk-based approach to aspirin therapy decision-making in older adults.

6
Patterns of deprescribing after an emergency department visit due to adverse drug events among concomitant users of antithrombotics and other medications

Chi, E.; Soliman, A.; Hunold, K. M.; Chiang, C.-W.; Unroe, K. T.; Nechi, R. N.; Caterino, J. M.; Li, L.; Zhang, P.; Donneyong, M.

2026-07-22 epidemiology 10.64898/2026.07.20.26358502 medRxiv
Top 0.1%
11.1%
Show abstract

Objectives: This study aimed to examine patterns of deprescribing practices that are implemented after an emergency department (ED) visit due to antithrombotic-induced gastrointestinal (GI) bleeds among patients who concomitantly use antithrombotic and other medications. Methods: A retrospective cohort study design was used to identify a cohort of patients who visited an ED from the MarketScan claims database (2016-2023). Concomitant use of antithrombotics and other medications was assessed in the 30 days prior to the ED date (index date). Those who presented with GI bleeds on their ED visit were classified as exposed vs those without GI bleeds (unexposed). Antithrombotic deprescribing-defined as discontinuation ([≥]45 days gap between refills), switching, or dose reduction-was assessed from the index date through to end of data. Inverse probability of treatment (IPT)-weighted logistic regression models were used to compare the odds of deprescribing between the exposed vs. unexposed groups. Results: Of the 375,510 total concomitant users of antithrombotics and other agents, 9,145 had a GI bleed vs. not (366,365). The odds of antithrombotic deprescribing was significantly higher among the exposed vs unexposed, (odds ratio [OR], 1.39; 95% confidence interval [CI], 1.33, 1.46). There was no significant difference in the discontinuation of concomitant medications overall. However, among patients who continued antithrombotic use after the ED visit, the discontinuation of concomitant medications was relatively higher among the exposed (OR, 1.11; 95% CI, 1.05, 1.16). Conclusions: Antithrombotic agents were more likely to be deprescribed after an ED visit among patients who concomitantly used antithrombotics and other medications.

7
Systematic Data Fitness Assessment Improves Validity and Replicability of Research Using Real-World Data

Razzaghi, H.; Wieand, K.; Pinkney, A.; Bailey, C.

2026-08-10 epidemiology 10.64898/2026.08.05.26359818 medRxiv
Top 0.1%
10.1%
Show abstract

Research replication is essential to build trust in evidence produced from real-world data. However, methods for conducting and reporting these studies are lacking, particularly related to data quality and fitness assessments. We replicated a single-center study from Children's Hospital of Atlanta in a multi-institutional learning network (PEDSnet) to evaluate the long-term effects of hydroxyurea in children with severe sickle cell disease (SS/S{beta}0 genotype). An AS-IS arm applied the original study's criteria with no major data quality adjustments, while a Data Fitness Enhanced (DFE) arm used systematic data fitness assessment to inform adjustments to cohort inclusion criteria and variable definitions; both arms then replicated the original study's primary analyses. Data quality checks in the DFE arm refined cohort criteria and improved hydroxyurea capture, drug era computation, and hematology specialist mapping. The DFE cohort produced average treatment effects with higher face validity and greater concordance with the original study (e.g., change in ED visits: -0.44 (CI -0.60, -0.26) versus -0.36 (CI -0.57, -0.16) in the original study) than the AS-IS cohort (-0.08 (CI -0.26, 0.09)), which yielded several implausible results. These findings show that superficially plausible cohort characteristics do not guarantee valid results without transparent, systematic data fitness assessment.

8
Identifying anaphylaxis using weakly-supervised prediction models and natural language processing

Williamson, B. D.; Cronkite, D. J.; Yu, O.; Ramaprasan, A.; Fuller, S.; Covey, J.; Kiniry, E.; Park, D.; Winter, R.; Whitaker, J.; McLemore, M. F.; Wittayanukorn, S.; Stojanovic, D.; Zhao, Y.; Dutcher, S.; Carrell, D. S.; Jackson, L. A.; Nelson, J. C.; Smith, J. C.

2026-06-17 epidemiology 10.64898/2026.06.09.26355005 medRxiv
Top 0.1%
9.2%
Show abstract

Objectives Scalable computable phenotyping algorithms are critical for conducting high-throughput disease-outcome research in large, distributed-data electronic health record (EHR) and claims data settings. We developed and evaluated a claims- and EHR-based computable phenotyping algorithm for anaphylaxis, a rare acute condition that is challenging to accurately identify using claims data alone. Materials and Methods Potential anaphylaxis events came from two healthcare systems (Kaiser Permanente Washington [KPWA] and Vanderbilt University Medical Center [VUMC]). We engineered features from clinical text using automated natural language processing (NLP) methods. We then developed a phenotyping algorithm using four NLP- and diagnosis code-based silver labels (proxies for the gold-standard labels). Gold-standard abstracted outcomes were used to evaluate algorithm performance. Results The largest area under the receiver operating characteristic curve (AUC) was 0.931 for an NLP-based silver-label model at KPWA. Depending on the model and healthcare system site, positive predictive value (PPV) and sensitivity at the threshold of predicted probability that maximized F1 score ranged from 0.52 to 0.77 (PPV) and 0.78 to 1 (sensitivity). Discussion NLP-based silver-label models had large AUC at KPWA but not at VUMC. This may be because clinical text at KPWA is only available for outpatient encounters and secure messaging. High sensitivity for identifying anaphylaxis can be obtained using our best-performing models. Conclusion The best-performing models had better PPV and sensitivity tradeoffs than prior bespoke anaphylaxis models with costly, manually curated features. The simplicity of the approach compared to traditional phenotyping methods allows it to be deployed easily at multiple health care systems.

9
Automating clinical trial outcome identification and misreporting detection using RegCheck

Cummins, J.; Drysdale, H.; Elson, M.; Hussey, I.; Goldacre, B.; DeVito, N. J.

2026-08-02 epidemiology 10.64898/2026.07.30.26358885 medRxiv
Top 0.1%
8.9%
Show abstract

Objective To evaluate the accuracy and cost of RegCheck, an automated large language model (LLM)-based workflow, for identifying clinical trial outcomes and detecting outcome misreporting by comparing its outputs with manual assessments from the COMPare Trials project. Design Validation study. Setting Sixty-two clinical trials originally assessed in the COMPare Trials project, sampled from five high impact general medical journals. Participants Published clinical trial reports and their corresponding prespecified registrations and/or protocols. Main outcome measures Four prespecified research questions were examined. RQ1 assessed outcome extraction recall relative to COMPare. RQ2 assessed accuracy of outcome classification as primary, secondary, or non-prespecified. RQ3 assessed accuracy of misreporting detection relative to COMPare, with additional manual adjudication of discrepancies between RegCheck and COMPare. RQ4 assessed the average per-paper cost of running the automated workflow. Results Across the validation papers, RegCheck achieved 91.2% outcome extraction recall relative to COMPare, and 83.6% outcome classification accuracy. For detection of outcome misreporting, RegCheck's overall accuracy was 85.6%. However, after resolving discrepancies with the original human judgements (which frequently favoured RegCheck's judgement), revised accuracy for outcome misreporting detection was 94.8%. The mean cost of running the workflow was 5.94 USD per paper. Conclusions RegCheck achieved high overall performance with a rigorous manual benchmark for identifying prespecified and reported outcomes in clinical trials, and detecting outcome misreporting, while operating at very low marginal cost. Adjudication of discrepant judgements suggested that RegCheck frequently identified valid issues not captured in the reference standard. Automated outcome checking may offer a scalable way to support editors, peer reviewers, and authors in detecting outcome switching and improving trial reporting.

10
Escalate or Switch? Treating the Post-Titration GLP-1 Non-Responder: A Target Trial Emulation With Dose-Equivalence Reclassification

Erly, B.; Raja, S.

2026-07-16 epidemiology 10.64898/2026.07.14.26357491 medRxiv
Top 0.1%
7.3%
Show abstract

Background. When a GLP-1 patient stops responding, should the clinician push the dose or change the drug? Observational answers conflate two distinct sources of confounding. Most early-week "escalations" in real-world data are FDA-mandated titration steps rather than deliberate clinical decisions, and patients who deviate do so for reasons we cannot observe. Semaglutide and tirzepatide are also not equivalent milligram-for-milligram, so naive class-switch comparisons mix mechanism and dose. We resolve both by restricting to post-titration patients and reclassifying treatments under the Whitley 2023 dose-equivalence framework. Methods. From 68,969 telehealth GLP-1 patients we built a post-titration cohort. Each patient's index time is the day they completed at least four weeks at therapeutic dose (Whitley tier 3 or higher: semaglutide 1.0 mg or tirzepatide 5 mg). Confirmed slow response is less than 5% total weight loss at the index, consistent with FDA weight-management drug-development guidance and AACE/ACE criteria. We compared four post-index strategies against continuing the current regimen: within-class dose escalation, equipotent class switch (a Whitley tier change of 1 or fewer), and class switch with potency increase. Direction-specific analyses split switches into semaglutide-to-tirzepatide and tirzepatide-to-semaglutide arms. Outcomes were percent weight loss at 12 and 24 weeks post-index. We estimated effects six ways: propensity-score matching; IPTW with linear and gradient-boosted propensities; the g-formula with linear and gradient-boosted outcome models; and AIPW, the doubly-robust estimator we use as the tiebreaker. Two-layer inverse probability of censoring weighting addressed strategy adherence and outcome ascertainment. We computed E-values, ran a negative-control specification, and stratified by tolerability. Results. The post-titration cohort comprised 24,876 confirmed slow responders. Within-class dose escalation produced a small consistent benefit at 24 weeks: AIPW +0.64 pp (95% CI +0.16 to +1.12), with five non-AIPW estimators ranging +0.47 to +0.76 pp. The continue arm itself lost an additional 8.07 pp over the same window (96% continued to lose), so escalation is a marginal addition to a substantial natural slope, not a rescue. Equipotent class switching from semaglutide to tirzepatide was inconclusive: linear and matching estimators ranged +1.26 to +1.80 pp, but AIPW was -0.33 pp (95% CI -1.27 to +0.60) with only 90 treated patients and limited propensity-score overlap. Class switch with simultaneous potency increase (sema to tirz) gave AIPW +0.65 pp at 12 weeks (95% CI +0.37 to +0.92, n = 80). A negative-control specification yielded ATE -0.13 pp, indicating the pipeline did not generate spurious signal. A held-out-fold prognostic-threshold sensitivity gave a null effect (+0.04 pp), correcting an earlier circular +1.06 pp estimate. Conclusions. Among confirmed post-titration slow responders, within-class dose escalation adds approximately 0.6 percentage points at 24 weeks on top of an 8 percentage point natural slope, consistently across six estimators including doubly-robust inference. This headline effect is small and not robust to modest unmeasured confounding (E-value 1.27) or to MNAR-style outcome attrition (tipping point delta approximately 1.2 pp), so it should be read as hypothesis-generating rather than practice-changing. Class-switching evidence is inconclusive; linear-estimator results suggesting benefit did not survive doubly-robust estimation in small treated samples with limited propensity overlap. The dose-ladder framework, with phase-specific evidence grading, is hypothesis-generating and insufficient on its own to change practice.

11
Same Result, Different Price: Compounded versus Branded Tirzepatide

Erly, B.; Raja, S.

2026-07-16 pharmacology and therapeutics 10.64898/2026.07.14.26357505 medRxiv
Top 0.1%
6.8%
Show abstract

Background. Compounded tirzepatide is prescribed at scale as a cheaper substitute for branded Mounjaro and Zepbound, yet the cost case is almost always built by setting one compounded price against one branded list price. That framing ignores the question that actually decides the answer: cheaper than which branded price the patient can reach. Branded tirzepatide is now sold at sharply different tiers, namely insurance copay (often $25-$150/month), LillyDirect Self Pay ($299-$449/month), and retail cash price ($1,000-$1,200/month). Whether compounded saves money turns entirely on which of these a given patient faces. A second open question is whether the two formulations even produce comparable effectiveness, since observed differences may reflect selection on insurance, baseline characteristics, and adherence rather than the drug. Methods. We conducted a retrospective cohort study of tirzepatide users in the Mochi Health telehealth program, classified by formulation from their refills as branded-only (Mounjaro/Zepbound; 6,238), compounded-only (71,683), or switchers (4,996); switchers were excluded from the formulation contrast. Among single-formulation patients with a documented six-month weight observation, the analytic cohort was 7,271 (869 branded, 6,402 compounded). The primary outcome was six-month percent body weight loss; the secondary outcome was >=10% response. We used 1:1 nearest-neighbor propensity-score matching (0.25 SD caliper) on baseline covariates only - age, sex, baseline BMI, baseline weight, comorbid diabetes, hypertension, dyslipidemia, prior bariatric surgery, and self-reported insurance coverage - deliberately excluding post-treatment variables such as adherence and time in program, which are mediators of the formulation effect. We pre-specified an equivalence margin of +/-2 percentage points on mean loss and tested equivalence with two one-sided tests (TOST). A directed acyclic graph (DAG) makes the identifying assumptions explicit; metformin use could not be reliably ascertained and is treated as an unmeasured confounder. The cost comparison reports the savings or premium of compounded versus branded under five branded price scenarios: retail list, LillyDirect Self Pay (two dose tiers), and insurance copay (typical and low end). It is a cost comparison (cost-minimization under demonstrated similar effectiveness), not a formal cost-effectiveness analysis: we computed no ICER, QALY, or discounting. Results. Branded and compounded patients had similar outcomes even before adjustment (mean loss 11.7% vs 11.5%; >=10% response 60.9% vs 58.8%). The largest baseline difference between the groups was insurance coverage (branded patients far more likely insured; standardized mean difference 0.67), which matching balanced to 0.01. After 1:1 matching (718 pairs, all |SMD| < 0.04), mean loss was 11.4% vs 11.4% (difference +0.08 pp, 95% CI -0.70 to +0.80) and >=10% response 59.3% vs 57.2% (difference +2.1 pp, 95% CI -3.1 to +7.1). The two formulations were statistically equivalent within the pre-specified +/-2 pp margin (TOST p < 0.001). Cost depends on the branded scenario: compounded saves $6,000 over six months versus retail list price, $1,494 versus LillyDirect maintenance-dose (5-15 mg) Self Pay, and $594 over a low-dose (2.5 mg) LillyDirect prescription, while it costs $300 more than branded under a typical insurance copay ($150/month) and is more expensive still at lower copays (savings turn negative below $200/month). Conclusions. Branded and compounded tirzepatide were statistically equivalent in six-month effectiveness within a pre-specified +/-2 pp margin, so the choice between them is essentially a cost decision - and that cost advantage is real but conditional on the branded price the patient can access. It is large against retail list price and shrinks to zero or reverses against LillyDirect Self Pay or a low insurance copay. Whether compounded is the lower-cost choice for an individual patient is, therefore, a question about which price tier that patient faces.

12
DrugSet: A validated R Shiny application for reproducible drug codelist construction from ATC classification to CPRD Aurum prodcodes

Hoxhaj, V.; Fry, C.; Morris, D.; Aurelius, T.; Martin, S.; Sturkenboom, M.; Andaur Navarro, C.

2026-07-13 epidemiology 10.64898/2026.07.08.26357534 medRxiv
Top 0.1%
6.7%
Show abstract

Objectives. To present DrugSet, a validated R Shiny application supporting the construction medicinal products codelists based on the Anatomical Therapeutic Chemical (ATC) system and their mapping to Clinical Practice Research Datalink (CPRD) Aurum prodcodes within a single interactive workflow. Materials and Methods. DrugSet comprises four modules: data preparation, ATC-based hierarchical code selection, string-based CPRD Aurum prodcodes mapping, and codelist export. Validation was conducted against World Health Organization (WHO) ATC reference codelists and manually curated prodcodes mappings across three drug classes: metformin, beta-blocking agents, and topical salicylic acid. Sensitivity, specificity, and Positive Predictive Values (PPV) were calculated for ATC codelist generation. Agreement proportions (overlapping against total identified codes) were calculated for prodcodes mapping. Time needed for codelist construction using DrugSet was recorded and compared to manual approaches. Results. DrugSet ATC codelist generation against WHO manual reference achieved 100% sensitivity, specificity, and PPV across all medicinal products. Prodcodes mapping agreement ranged from 89.2% to 98.3% with discrepancies due to missing data in the prodcodes input vocabulary. DrugSet completed codelist construction in 9 minutes compared to 3 hours and 10 minutes manually, across all medicinal products classes. Discussion. DrugSet provides a unified workflow that runs directly on ATC and source CPRD Aurum vocabulary files. The reduction in codelist construction time and export of the generated codelists supports reproducibility in pharmacoepidemiologic studies where codelist creation can represent a significant proportion of study setup time. Conclusion. DrugSet is an open-source, validated tool that improves accuracy, and efficiency of codelist construction for medicinal products based on ATC codes towards CPRD Aurum prodcodes.

13
A Renal Safety Checkpoint for Early High Intensity Statin Therapy in Critically Ill Patients With Acute Coronary Syndrome: A Multidatabase Target Trial Emulation

Huang, K.; Zheng, X.; Liu, J.; Wu, C.; Sun, H.

2026-08-07 cardiovascular medicine 10.64898/2026.08.05.26359828 medRxiv
Top 0.1%
6.6%
Show abstract

Background: High intensity statins are foundational after acute coronary syndrome (ACS), yet intensive care unit prescribing occurs while renal reserve, perfusion, and interacting therapies are changing. We tested a renal safety checkpoint integrating kidney status, hemodynamic instability, and drug interaction burden to identify when statin intensity may become nonexchangeable. Methods: We emulated an active-comparator target trial across MIMIC-IV, eICU, and MIMIC-III. Critically ill adults with ACS, acute myocardial infarction, or percutaneous coronary intervention who received high- or moderate-intensity statins within 24 hours were included. The primary outcome was 7-day KDIGO stage 2 or 3 acute kidney injury or incident renal replacement therapy. Eligibility, time zero, treatment assignment, and follow-up were aligned. Database-specific propensity scores, overlap weighting, and standardization addressed confounding and treatment overlap. Safety domains, longitudinal analyses, bootstrap resampling, source omission, and endpoint sensitivities assessed robustness. Results: Among 5,178 patients, 761 developed the primary outcome, including 223 who initiated renal replacement therapy. Standardized risks were 17.40% with high-intensity therapy and 15.01% with moderate-intensity therapy (risk difference, 2.39 percentage points [95% confidence interval (CI), -0.23 to 5.05]; risk ratio, 1.16 [95% CI, 0.99 to 1.39]). Risk separation was greatest with high hemodynamic instability (5.78 percentage points [95% CI, 1.56 to 9.74]) and high drug-interaction burden (6.24 percentage points [95% CI, -0.44 to 12.19]). Renal replacement therapy showed a 1.33-point risk difference (95% CI, 0.18 to 2.67). Conclusions: This study moves statin safety assessment beyond fixed dose label or isolated creatinine measurement. The findings support a clinically actionable monitoring strategy in which early statin intensity is reassessed against evolving perfusion, kidney status, and interaction burden. This approach preserves intensive lipid lowering for physiologically suitable patients while identifying a high risk window in which temporary moderation.

14
Assessing Equity and Representativeness in Randomised Controlled Trials: A Feasibility Study

Oparah, C.; O'Keefe, H.; Agbeleye, O.; Nesworthy, J.; Norman, G.; Kunonga, T. P.

2026-07-06 epidemiology 10.64898/2026.06.25.26356548 medRxiv
Top 0.1%
5.6%
Show abstract

Clinical trials often enrol populations that differ from those who ultimately receive the interventions, raising concerns about external validity and health equity. Trial registries could provide an early opportunity to assess representativeness, but it is unclear whether registry data contain sufficient information to enable such assessments. This study evaluated the feasibility of using registry data to assess representativeness in Phase II and III pharmacological randomised controlled trials. A search of ClinicalTrials.gov from December 2024 to January 2025 identified trials with results posted after 1 January 2023 across cardiovascular disease (CVD) excluding stroke, diabetes mellitus, and selected mental health disorders. Of 1,328 records screened, 98 trials met inclusion criteria (51 Phase III, 47 Phase II). Reporting completeness was variable, particularly in Phase II studies. CVD and diabetes trials predominantly included middle-aged to older adults, while mental health trials recruited mainly individuals aged 36 to 50 years. Across CVD and mental health trials, participants were largely male. Reporting of BMI, contraception, and comorbidity criteria was inconsistent, though available data suggested these factors influenced sample composition. Fewer than 10% of trials reported equity-relevant characteristics beyond age and sex, and none addressed intersectionality. Assessing equity using registry data is feasible but constrained by incomplete and inconsistent reporting.

15
Feasibility of using automatically extracted routine clinical data in a respiratory cohort study: The SPHN-SPAC demonstrator project.

Romero, F.; Sasaki, M.; Mallet, M. C.; Pedersen, E. S. L.; Leuenberger, L. M.; Makhoul, R.; Bovermann, X.; Hartung, A.; Latzin, P.; Kissling, S.; Moeller, A.; Treis, A.; Regamey, N.; Belle, F. N.; Kuehni, C. E.

2026-07-16 epidemiology 10.64898/2026.07.14.26357927 medRxiv
Top 0.1%
4.9%
Show abstract

Objectives To assess the feasibility of using clinical data automatically extracted via the Swiss Personalized Health Network (SPHN) to complement or replace manually abstracted clinical data in the Swiss Paediatric Airway Cohort (SPAC). Materials and Methods We studied 1,075 SPAC participants enrolled between 2017-2023 at two Swiss children's hospitals. Clinical data were extracted from electronic health records via SPHN in Resource Description Framework format, transformed into visit-centered datasets, and compared with manually abstracted SPAC clinical data and parent-reported emergency department (ED) visits and hospitalizations from follow-up questionnaires. We assessed feasibility by identifying challenges in acquiring data and evaluated data quantity, completeness, and agreement between datasets. Results We obtained analysis-ready SPHN-derived datasets from two hospitals after 24 months. SPHN-derived data captured more pneumology outpatient visits than manual abstraction (Hospital A: 1,963 vs 1,049; Hospital B: 2,343 vs 1,010) and identified clinical events among children without follow-up questionnaires. Completeness of variables varied across hospitals and encounters, reflecting differences in local clinical documentation practices. SPHN-derived and manually abstracted data showed high agreement for structured clinical variables, including spirometry measurements (concordance correlation coefficient >0.99). Self-reported and SPHN-derived ED visits and hospitalizations showed high absolute agreement but moderate concordance. Discussion and Conclusion Automated extraction of routine clinical data increased the completeness of longitudinal information compared with manual abstraction, suggesting that SPHN-derived data can complement manual data collection in cohort studies. Broader use remains limited by heterogeneous clinical documentation practices and the substantial effort required to harmonize and transform extracted data into analysis-ready research datasets.

16
Publication Bias in Abstracts Presented at the American Diabetes Association Scientific Sessions: A Retrospective Cohort Study

Pinedo-Torres, I.; Taype-Rondan, A.; Vera-Luza, A. A.; Zegarra-Lizana, P. A.; Rojas-Vilca, J. L.; Yovera-Aldana, M.

2026-08-31 epidemiology 10.64898/2026.08.26.26361486 medRxiv
Top 0.1%
4.2%
Show abstract

Objective. To determine the publication rate of abstracts presented at the American Diabetes Association Scientific Sessions and to evaluate the association between statistical significance of study results and subsequent publication. Research Design and Methods. We conducted a retrospective cohort study of abstracts presented at the 2018 American Diabetes Association Scientific Sessions. The primary exposure was study result category (statistically significant vs. non-statistically significant findings), and the primary outcome was publication in an indexed journal within 5 years after conference presentation. Publication status was determined through PubMed/MEDLINE and Scopus searches. Adjusted relative risks (RRs) and 95% CIs were estimated using generalized linear models with Poisson distribution and robust variance. Results. Among 541 included abstracts, 321 (59.3%) were subsequently published in indexed journals. Abstracts reporting statistically significant findings had a higher publication rate than those reporting non-statistically significant findings (61.9% vs. 42.3%; p=0.002). In the adjusted analysis, abstracts with non-statistically significant findings had a lower likelihood of publication compared with those reporting statistically significant findings (adjusted RR 0.71 [95% CI 0.55-0.93]; p=0.013). Conclusions. Approximately four in ten abstracts presented at the ADA Scientific Sessions were not published within 5 years. Abstracts reporting non-statistically significant findings had a lower likelihood of subsequent publication, suggesting persistent publication bias in diabetology research. Future initiatives promoting the interpretation of effect estimates, confidence intervals and clinical relevance, rather than statistical significance alone, may help reduce selective dissemination of evidence

17
ICD-10 Code Ambiguity Obscures Treatment-Eligible Adults with Spinal Muscular Atrophy: A Single-Center Chart Review and Patient Outreach Study

Holly, G.; Bean, B.; Beshay, H.; Edwards, G.; Streicher, N. S.

2026-06-15 neurology 10.64898/2026.06.07.26355122 medRxiv
Top 0.1%
4.1%
Show abstract

Background. Three disease-modifying therapies (DMTs) for spinal muscular atrophy (SMA) have been approved since 2016, yet many adults remain untreated. Identifying them depends on ICD-10 codes that capture SMA but do not reliably distinguish it from other related conditions. We examined, in one U.S. health system, both patients' engagement with therapy and the accuracy of the codes used to find them. Methods. We conducted a retrospective chart review of adults in an academic health system identified by SMA-associated ICD-10 codes, with manual adjudication of diagnosis and DMT status. Confirmed SMA-positive, DMT-naive patients were invited to a structured telephone interview on treatment awareness and barriers. Results. Of 60 charts, 22 (36.7%; 95% CI 25.6-49.3%) were appropriately coded for SMA or a related disorder; only 16 (26.7%) had molecularly confirmed SMA. The other 38 (63.3%) were miscoded, spanning spinal and bulbar muscular atrophy, asymptomatic carriers, prenatal screening, and conditions unrelated to SMA. Ten of the 16 confirmed patients (62.5%) were DMT-naive; one was interviewed, one declined, and eight could not be reached. The non-response is itself a finding: the patients least visible to administrative data are the hardest to reach. Conclusions. ICD-10 ambiguity is a barrier to treatment access in adult SMA, as is loss to follow-up. We make two recommendations: continuous documentation-coding alignment that uses natural language processing to verify the genetic precondition, and type-specific SMA codes (subcodes for Types 0-4) anchored on molecular SMN1 confirmation. Together these would support cohort identification, outreach, and evidence generation without adding to clinician burden.

18
A Simulation Study Comparing Multiple Imputation and Complete Case Analysis for Handling Missing Preschool Body Mass Index

Savu, A.; Dover, D. C.; Hajihosseini, M.; Gaudet, L. A.; Kaul, P.

2026-08-14 epidemiology 10.64898/2026.08.13.26360115 medRxiv
Top 0.1%
4.1%
Show abstract

Background and Objective. Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with. Using real-world data, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age, where missing values were present in as much as 30% of our sample. Methods. This study utilized a cohort of 130,424 children with complete preschool-age body mass index (BMI) measurements from the Calgary and Edmonton health regions of Alberta, Canada. In the complete BMI data, we introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). To handle the missing data created, we employed complete-case and multiple-imputation methods. Maternal glucose levels during pregnancy were categorized into five groups and its association with child excess weight at pre-school age was determined based on a logistic regression model using the full observed data (yielding true values), observed data that was not deleted (complete-case estimates), and imputed data (multiple-imputation estimates). The accuracy of complete-case and multiple-imputation estimates were evaluated against the true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift. Results. Under MCAR and MAR, multiple-imputation generally outperformed complete-case, yielding smaller absolute and relative bias. Both methods achieved high significance ([&ge;] 0.96) for most effects. Mean squared errors for multiple-imputation and complete-case were similar missing completely at random, missing at random, and coverage was consistently high ([&ge;] 0.99). Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis using pattern-mixture models indicated performance varied by specific effect. Conclusions. Under MCAR and MAR, multiple-imputation introduced higher bias but demonstrated superior overall performance based on mean squared error and restored statistical power. Conversely, both methods failed under MNAR, where pattern-mixture modeling sensitivity analyses revealed highly variable, effect-specific performance due to unverifiable shift assumptions. When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.

19
First-Line Opioids and Short-Term All-Cause Emergency Department Return After Headache Visits: A Two-Center Comparative Cohort Study

Gorenshtein, A.; Adiniaev, Y.; Liba, T.; Klang, E.; Daniel, O.

2026-07-17 neurology 10.64898/2026.07.16.26358169 medRxiv
Top 0.1%
3.5%
Show abstract

Objective: To compare first-line emergency department (ED) treatment classes for acute headache on short-term all-cause ED return and index admission across two independent health systems. Background: ED trials of acute headache treatment are judged on in-ED pain relief, a documented endpoint that is recorded incompletely and shifts with the scoring rule, and is a weak surrogate for what happens after discharge. All-cause ED return after an index headache visit (any subsequent ED encounter within the window) has not been used to compare first-line treatments at scale, and society guidance favors dopamine-receptor antagonists while recommending against routine opioids. Methods: Retrospective two-center cohort of adults treated for headache in the ED, using MIMIC-IV-ED (Beth Israel Deaconess Medical Center, 2011-2019) and MC-MED (Stanford, 2020-2022). The first-line class was the earliest qualifying acute agent. The primary contrast was opioids versus dopamine-receptor antagonists (the guideline-preferred class). Outcomes were 72-hour and 7-day all-cause ED return (among discharged patients; any subsequent ED encounter within the window) and index hospital admission. Confounding by indication was addressed with propensity overlap weighting; associations are reported as adjusted risk ratios (RRs) with bootstrap 95% CIs and E-values. Estimates were pooled with a site term and examined per site. Results: Among 13,285 treated adults (10,799 MIMIC-IV-ED; 2,486 MC-MED), opioid recipients were older and higher-acuity than dopamine-antagonist recipients (index admission 38.1% vs 16.4%). In the MIMIC-IV-ED discharged primary-contrast population, overlap weighting reduced the maximum standardized mean difference from 0.35 to 0.002; pooled and site-specific balance diagnostics are provided in the Supplement. First-line opioids remained associated with a higher 72-hour all-cause ED return (6.8% vs 3.8%; adjusted RR 1.79; 95% CI 1.31 to 2.33), 7-day return (10.7% vs 6.6%; RR 1.62; 95% CI 1.28 to 1.98), and index admission (RR 2.32; 95% CI 2.11 to 2.58, consistent with strong residual severity differences in patients selected for opioids). The direction of association was concordant across both health systems, although MC-MED return estimates were imprecise given the smaller opioid-treated discharged sample. In MIMIC-IV-ED, the cumulative all-cause return incidence by treatment class separated by day 3 and persisted through 30 days. The direction was consistent, though attenuated and no longer statistically significant, when the outcome was restricted to a headache-specific return (72-hour RR 1.31; 95% CI 0.91 to 1.88); the direction persisted for the composite of admission or 72-hour return, which does not condition on discharge but is influenced by the more confounded admission component (RR 2.16; 95% CI 1.98 to 2.39). Conclusion: Across two health systems, first-line opioid treatment for ED headache was associated with higher all-cause short-term ED return among discharged patients and higher index admission than dopamine antagonists. These observational associations reflect downstream all-cause ED utilization after an index headache visit rather than confirmed headache recurrence or treatment failure; they are consistent with guideline-concordant, opioid-sparing first-line treatment and warrant prospective confirmation. Plain Language Summary: Emergency departments treat headaches with several different medicines, but the usual way of judging which works, the pain score recorded during the visit, is often missing or inconsistent. Using two large hospital systems and a clearer outcome, whether patients came back to the emergency department for any reason, we found that patients first treated with opioids returned within 72 hours about 1.8 times as often as those given the guideline-preferred dopamine-blocking medicines and were admitted more than twice as often. These patterns pointed the same direction in both hospital systems after adjustment for the measured differences available in both databases. Because this was an observational comparison and returns were counted for any reason, the findings are consistent with using guideline-preferred non-opioid medicines first, rather than proof that opioids worsen headache.

20
Death in People with Down syndrome: Mortality statistics and novel predictors in US Medicaid and Medicare enrolled adults.

Tewolde, S.; Rosellini, A. J.; Michals, A.; Skotko, B. G.; Fortea, J.; Khor, B.; Handelman, S.; Rubenstein, E.

2026-07-20 epidemiology 10.64898/2026.07.17.26358090 medRxiv
Top 0.1%
3.4%
Show abstract

People with Down syndrome have higher age-specific mortality rates compared to the general population as well as peers with other intellectual and developmental disabilities. While a large proportion of mortality is attributable to Alzheimers disease, many die prior to Alzheimers diagnosis and some live to old ages, dying without Alzheimers. Our objectives were to use 11 years of Medicaid and Medicare data to describe characteristics and factors related to death in adults with Down syndrome and use machine learning to identify which conditions most strongly predict death in the full population and stratified by age. We identified death using Center for Medicare and Medicaid Systems reported date of death health conditions using ICD 9 and 10 codes. We used a case-control design with risk set sampling to have that controls to mimic the distribution of times of incident Alzheimers disease. We trained gradient boosted trees to identify strongest predictors. Our cohort included 137,293 adults with Down syndrome. Among those, 30,894 (22.5%) died during the study period. Mean age at death among those who died was 55 years (SD=10). Mean age of death in those with Alzheimers disease was 59 (SD=7) and those without was 52 (SD=12). The most influential predictors of mortality were any claim for dementia, any claim for pneumonia, re-occurring claim for cardiovascular disease three years before index death, and any claim for heart failure and epilepsy. Our results align with previous clinical work and highlight intervenable areas to reduce mortality in the Down syndrome population.